Skip to content

slurm: batch the queries and parallelize the writes in the GPU power/clock helpers - #1393

Open
100milliongold wants to merge 2 commits into
NVIDIA:masterfrom
xiilab:perf/exclusive-gpu-parallel
Open

slurm: batch the queries and parallelize the writes in the GPU power/clock helpers#1393
100milliongold wants to merge 2 commits into
NVIDIA:masterfrom
xiilab:perf/exclusive-gpu-parallel

Conversation

@100milliongold

Copy link
Copy Markdown
Contributor

Problem

set_gpu_power_levels.sh and set_gpu_clocks.sh call nvidia-smi once per GPU to read the target value and once more to apply it, all serially. Both nvidia-smi -pl and nvidia-smi -ac take roughly a second per GPU, so on an 8-GPU node the two helpers together add about 8 s to the prolog of every job for which 50-exclusive-gpu runs. srun surfaces it as:

srun: Prolog hung on node <node>

Observed on DGX OS 7.5.0 (8× B300), Slurm 26.05.1.

Fix

  • Read the values for all GPUs in a single --query-gpu call (the per-GPU -i loop was only needed because the value was read one at a time).
  • Apply them in parallel, collecting each child's exit status so a failure on any GPU still fails the script.

The same values are written to the same GPUs; only the number of nvidia-smi invocations and their concurrency change. The default branch of set_gpu_clocks.sh already operated on all GPUs at once and is untouched.

Verification

Not yet timed with the patched scripts on hardware — the system where this was found has 50-exclusive-gpu removed from prolog.d (an 8-GPU node shared between jobs should not have every job reset limits and clocks on all GPUs). The serial cost is reproducible there: prolog took 6–8 s per job while the script was in place. Marked as draft for that reason; happy to run a timed before/after if that would help.

Related

The reason every job ran 50-exclusive-gpu in the first place is a separate defect in the exclusive-job detection, addressed in #1391.

…clock helpers

set_gpu_power_levels.sh and set_gpu_clocks.sh called nvidia-smi once per
GPU to read the target value and once more to apply it, all serially. Both
"nvidia-smi -pl" and "nvidia-smi -ac" take roughly a second per GPU, so on
an 8-GPU node the two helpers together add about 8 s to the prolog of every
job that 50-exclusive-gpu runs for. srun reports this as:

  srun: Prolog hung on node <node>

Read the values for all GPUs in a single --query-gpu call, then apply them
in parallel and collect each child's exit status so a failure on any GPU
still fails the script.

Behaviour is otherwise unchanged: the same values are written to the same
GPUs. The "default" branch of set_gpu_clocks.sh already operated on all
GPUs at once and is untouched.

Observed on DGX OS 7.5.0 (8x B300), Slurm 26.05.1: prolog took 6-8 s per
job while 50-exclusive-gpu was running.

Signed-off-by: Jea-Eok-Kim <je.kim@xiilab.com>
@100milliongold
100milliongold marked this pull request as ready for review September 4, 2026 00:13

@dholt dholt left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The batched queries at set_gpu_power_levels.sh:21 and set_gpu_clocks.sh:13-14 run in process substitutions, whose failures are invisible to readarray and set -e. If a query fails with empty output, the loop is empty and the helper exits 0; partial output can update only some GPUs and also return success. Capture each query through a construct whose status can be checked, validate that all required per-GPU rows are present and aligned, and only then launch writes. Please cover empty, partial, and nonzero query results as required by the changed-path evidence.


Automated triage review (agent-generated on the maintainer's behalf; a human maintainer decides merges).

`readarray -t limits < <(nvidia-smi ...)` hides the query's exit status from
both readarray and `set -e`. The helpers therefore reported success in every
failure mode: an empty result made the write loop run zero times, a truncated
result configured only some of the GPUs, and a nonzero exit was not seen at all.

The query result now goes through a file so its status can be checked, the index
is selected alongside the values so a write targets the GPU nvidia-smi reported
instead of an array subscript, and every row is parsed and range-checked before
the first write is launched. The row count is compared against `nvidia-smi -L`,
so a partial result is rejected rather than silently applied.

set_gpu_clocks.sh additionally selected clocks.max.mem and clocks.max.sm in two
separate queries. If the two returned different row counts, `${maxMEM[$i]}` was
empty for the trailing GPUs and produced an `-ac ,1980` argument. Both values
now come from the same query, so they cannot drift out of alignment.

Verified on a DGX B300 (Ubuntu 24.04, bash 5.2) with an nvidia-smi stub that
honours --query-gpu and records writes instead of performing them. Identical
results for both helpers:

  case        before                        after
  ----------  ----------------------------  ---------------------
  ok          rc=0  8 writes                rc=0  8 writes
  empty       rc=0  0 writes                rc=1  0 writes
  partial     rc=0  3 writes (of 8 GPUs)    rc=1  0 writes
  fail        rc=0  0 writes                rc=1  0 writes
  nonnumeric  rc=0  8 writes ("N/A" passed) rc=1  0 writes

One limitation of the stub is worth stating: it records writes rather than
performing them, so the `nonnumeric` row shows 8 writes for the old code. On
real hardware `nvidia-smi -pl N/A` fails and `wait` would surface rc=1 — but
only after eight bad invocations. The new code rejects the row before the first
one.

The `ok` case is unchanged, so the parallel-write speedup this branch adds is
preserved.
@100milliongold

Copy link
Copy Markdown
Contributor Author

Thanks — the process-substitution point is correct, and the failure modes were worse than "invisible": all four of them returned success.

What changed (commit c95d3a2):

  • The query result now goes through a file so its exit status can be checked.
  • The index is selected alongside the values (--query-gpu=index,<field>), so a write targets the GPU that nvidia-smi reported instead of assuming the array subscript equals the GPU index.
  • Every row is parsed and range-checked, and the row count is compared against nvidia-smi -L, before the first write is launched. A partial result is rejected rather than silently applied.
  • set_gpu_clocks.sh selected clocks.max.mem and clocks.max.sm in two separate queries; if they returned different row counts, ${maxMEM[$i]} was empty for the trailing GPUs and produced an -ac ,1980 argument. Both values now come from the same query, so they cannot drift out of alignment.

Coverage of the cases you asked for, verified on a DGX B300 (Ubuntu 24.04, bash 5.2) with an nvidia-smi stub that honours --query-gpu and records writes instead of performing them. Identical results for both helpers:

case before after
ok (8 GPUs) rc=0, 8 writes rc=0, 8 writes
empty output rc=0, 0 writes rc=1, 0 writes
partial output (3 of 8) rc=0, 3 writes rc=1, 0 writes
nonzero exit rc=0, 0 writes rc=1, 0 writes
non-numeric value rc=0, 8 writes (N/A passed through) rc=1, 0 writes

One limitation of the stub is worth stating plainly: it records writes rather than performing them, so the non-numeric row shows 8 writes for the old code. On real hardware nvidia-smi -pl N/A fails and wait would surface rc=1 — but only after eight bad invocations. The new code rejects the row before the first one.

The ok case is unchanged, so the parallel-write speedup this branch adds is preserved.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants